AI Learning Series · Part 21

SOTA LLMs & Their MoE Architectures

Docs 02, 08 and 10 covered transformers, the memory wall and the optimization stack conceptually. This is the concrete survey: which architectures the frontier models actually use, and the engineering reasoning behind each number on the spec sheet.

Transformers
→
Optimization Stack
→
KV Cache Types
→
MoE Architectures
→
New Architectures

01 The Big Picture

Every frontier model spec sheet now reads like a split personality: "671B parameters" in the headline, "37B active parameters" in the footnote. That footnote is the whole story of modern LLM engineering.

Dense models wire every parameter into every token — capacity and per-token cost are the same number, growing together until serving becomes impossible (doc 08's memory wall). Mixture-of-Experts (MoE) breaks that coupling: store a huge parameter pool, but route each token through only a small, learned subset. The result — publicly reported across the current frontier — is models with dense-class latency and an order of magnitude more knowledge capacity.

This doc is the survey behind the concepts: the gate math, the load-balancing tricks, the capacity factor, the memory budget — and a model-by-model read of what Mixtral, DeepSeek-V3, GLM-4.5, Qwen3, Llama 4 and GPT-OSS actually shipped, as publicly reported.

🧭
Position in the series: doc 02 introduced MoE as one idea among many; doc 10 placed it in the serving stack. Here we open the box: the router, the experts, and the accounting that decides whether a deployment fits on your hardware. Attention-side memory tricks (MLA, KV types) live in doc 22; what might replace transformers entirely is doc 25.

02 What MoE Is — Precisely

A MoE layer replaces one feed-forward network (the FFN, which is ~2/3 of a transformer's parameters — doc 02) with N expert FFNs plus a router. For each token, the router scores all experts, picks the top-k, and blends only those outputs. Attention layers stay dense and shared by everyone.

Coarse MoE — Mixtral 8×7B

The design that made open MoE credible: 8 large experts (each ~7B params), top-2 routing. Total ≈ 47B parameters, but each token touches only ~13B — two experts plus the shared dense parts. Fine-grained it is not: with 8 experts and top-2 there are only C(8,2) = 28 possible expert combinations per layer.

Fine-grained MoE — the DeepSeek lineage

Instead of few large experts, use many small ones: DeepSeek-V3 (as publicly reported) splits its FFN into 256 routed experts, picks top-8, plus 1 shared expert per token. Combinatorially huge routing space → far more precise specialization per token, at the same active-parameter budget.

Why did fine-grained beat the 8-expert design? Combinations. With 8 experts top-2, a token's "sentence" of expert choices is 2 letters from an 8-letter alphabet. With 256 experts top-8, it's 8 letters from a 256-letter alphabet — the router can compose fine-grained capabilities (syntax + domain + format) instead of choosing between a few monolithic personalities. Same bytes read per token, much richer conditioning.

03 Why Sparsity Wins — Decoupling Capacity from Compute

The entire economic argument of MoE is one decoupling:

Dense model: capacity = cost. A dense 70B model stores 70B params and reads 70B params per token. To know more, it must cost more on every single token — even when a token only needs a sliver of that knowledge.
MoE model: capacity ≠ cost. DeepSeek-V3 stores 671B params but reads ~37B per token (publicly reported). Knowledge capacity scales with total params (weights in HBM); per-token compute and decode latency scale with active params (weights through the memory wall — doc 08). You grow the library without growing the librarian's walking route.
The dense-class anchor. 37B active lands in the "dense model you could actually serve" territory (~30–40B). That is the design target: a user-visible experience indistinguishable from a well-behaved dense model, backed by 671B of stored knowledge. Bigger active = better quality but slower and pricier per token; the frontier picks the point on that curve deliberately.
⚖️
The trade you accept: all 671B must still fit in memory on every serving replica, and routing adds load-balancing and all-to-all communication problems that dense models never have. MoE trades FLOPs for memory footprint and systems complexity. Section 07 does that accounting.

04 How It Works — One Token Through the Router

Step through a single token hitting one MoE layer. Watch the gate score every expert, keep only the top-k, and blend.

token x ROUTER W_g·x → softmax gate scores g(x) — one bar per expert E1 E2 E3 E4 E5 top-k = 2 kept → E2 (g=0.55), E4 (g=0.45) · E1/E3/E5 masked to 0 E1 · idle E2 · FIRES weight 0.55 E3 · idle E4 · FIRES weight 0.45 E5 · idle 3 of 5 experts untouched — their bytes never cross the bus y = Σ gᵢ·Eᵢ(x) 0.55·E2 + 0.45·E4 SHARED EXPERT always on, every token output = shared knowledge + routed specialists → next MoE layer / attention

Every MoE layer in the model runs this same ceremony, per token, per layer. The router is itself a tiny learned matrix (W_g) — its only job is to spend the compute budget wisely. Everything else in this doc is about what goes wrong when 30 billion tokens all want the same expert at once.

05 The Gate Math

Four formulas run the entire MoE economy. All symbols: T = tokens in batch, N = number of experts, k = top-k.

// 1. The gate — score every expert, keep top-k, renormalize g(x) = TopK( softmax(W_g · x), k ) // k nonzero weights, others = 0 y(x) = Σᵢ gᵢ(x) · Eᵢ(x) // blend ONLY the fired experts // 2. Expert capacity — how many tokens each expert may accept per batch C = ⌈ α · (T · k) / N ⌉ // α = capacity factor (~1.25 historically) // overflow → token DROPPED (passes through via residual connection) // 3. Load-balance auxiliary loss (Switch/GShard era) — added to training loss L_aux = α_lb · N · Σᵢ fᵢ · Pᵢ // fᵢ = fraction of tokens routed to i // Pᵢ = average gate prob mass given to i // 4. DeepSeek's aux-loss-free balancing — bias the SELECTION, not the gate value sᵢ = u·eᵢ + bᵢ // bᵢ used ONLY to pick top-k; gate value still softmax(u·eᵢ) bᵢ ← bᵢ + γ sign(overloadᵢ) // overloaded expert: bias down; hungry expert: bias up

Why the auxiliary loss was painful: L_aux pushes routing toward uniformity, but uniform routing is not what the model wants to learn — experts should specialize unevenly. The gradient conflict between "balance for hardware" and "specialize for quality" directly cost model performance. DeepSeek's reported fix decouples the two: the bias term bᵢ steers selection to keep experts busy while the actual gate values — what gets blended into the output — stay pure softmax of the true affinities. Balance becomes a control loop, not a gradient penalty.

Why capacity α ≈ 1.25: perfectly balanced routing would send exactly T·k/N tokens to each expert; the 1.25 headroom absorbs imbalance so a modestly popular expert doesn't overflow. It's a bargaining point: higher α wastes compute and buffer memory on idle slots; lower α drops more tokens (each dropped token's knowledge contribution for that layer is lost). Modern deployments without hard capacity limits instead penalize overflow softly — but the C formula is the vocabulary in every MoE paper.

Shared-expert regularization: the shared expert (always fired, every token) is told: "absorb the common denominator." Routing noise — trivia every token needs, grammar, generic transformations — no longer needs to be competed for in the top-k. That keeps routed experts free to specialize, and it's why the pattern (1 shared + k routed) recurs across DeepSeek-V3, GLM-4.5 and Llama 4 (as publicly reported).

06 Model-by-Model — The Frontier Spec Sheet

All figures below are publicly reported values as of writing; treat them as design landmarks, not contracts.

ModelParams total → activeExperts / top-kAttentionContextEngineering fit
Mixtral 8×7B 47B → 13B 8 large / top-2 GQA + sliding window 32K The existence proof: open MoE matching a much larger dense model at 13B-class latency. Coarse experts — 28 possible combinations per layer.
DeepSeek-V3 671B → 37B 256 fine-grained / top-8 + 1 shared MLA (KV compression — doc 22) 128K Frontier quality at ~dense-37B serving cost. Aux-loss-free balancing, multi-token prediction (MTP) head for speculative-style decoding.
GLM-4.5 355B → 32B 160 / top-8 + 1 shared (reported) GQA 128K Tuned explicitly for agentic coding + tool use; a smaller "Air" sibling (106B/12B) for cheaper serving.
Qwen3-235B-A22B 235B → 22B 128 / top-8 (reported) GQA + QK-Norm 128K Open-weight flagship with mature tooling; the "A22B" naming convention itself encodes total-vs-active.
Llama 4 (Scout / Maverick) 109B / 400B → 17B 16 / 128 routed + shared; alternating dense–MoE layers iRoPE (interleaved) for very long context up to 10M claimed (Scout) MoE only on some layers — a hybrid budget; native multimodal, huge-context ambitions.
GPT-OSS (120B / 20B) 117B / 21B → 5.1B / 3.6B 128 / top-4 (reported) attention sinks 128K Open-weights MoE small enough for a single consumer GPU; sinks stabilize attention over long generations.
📊
Read the pattern: active params cluster in the 13–37B band — the "servable dense model" zone — while total params climb 47B → 671B. The frontier is not racing to activate more; it's racing to store more while activating the same. GPT-OSS pushes the same logic down to 5B active for local use.

07 Engineering Takeaways — Serving the Beast

Memory footprint is driven by experts, not activations. You must hold every expert's weights on every replica regardless of how rarely some fire. DeepSeek-V3 at FP8: 671B × 1 byte ≈ 0.7 TB of weights alone. An 80 GB GPU holds ~12% of it. No single accelerator serves this — MoE is natively multi-GPU.
Expert parallelism (EP). The standard sharding: different experts live on different GPUs; tokens are dispatched across the network (all-to-all) each MoE layer, computed, and shipped back. EP load depends on routing balance — a hot expert becomes a hot GPU. This is why the load-balance math of section 05 is a systems concern, not just a training concern.
Roofline for tokens/sec (doc 08's rule, MoE edition). Decode is bandwidth-bound: token/s ≤ aggregate HBM bandwidth ÷ active bytes per token. 37B active @ FP8 ≈ 37 GB read per token per pass; an 8-GPU node with ~4–5 TB/s each gives a theoretical ceiling near ~1000 tok/s — before all-to-all traffic, KV reads, and overheads, which is why reported serving speeds sit far below it. Active params, not total, set your speed.
Compute per token stays small. Forward-pass FLOPs ≈ 2 × active params → ~74 GFLOPs per token for a 37B-active model. Compare capacity: 671B total. The ratio (~18×) is exactly the decoupling MoE sells.
✓ Do

Size your serving fleet by total params (memory) and budget speed by active params (bandwidth). Prefer models with a shared expert + aux-loss-free balancing for stable EP utilization. Watch expert-balance telemetry like you watch GPU utilization.

✗ Don't

Don't assume "671B model" means 671B-slow — or 13B cheap to host. Don't plan single-GPU serving for large MoE. Don't compare models on total params alone: a 235B-A22B and a 355B-32B are in the same latency class.

08 Mental Models

Hospital triage

A hospital employs hundreds of specialists (total params) but a patient only consults two or three per visit (active params). The payroll — memory footprint — is driven by the full staff; visit cost by the consultants seen. Triage (the router) must also prevent everyone queueing for the same cardiologist — that's load balancing.

A hospital's specialists are chosen by a human receptionist with common sense; the router learned its triage purely from gradient signal — and can develop irrational queueing habits the aux loss or bias terms must correct.
Library vs. librarian

Total params = the size of the library's collection; active params = how many shelves the librarian physically walks to per question. You can grow the collection indefinitely while keeping the walk short — but the building (HBM) must still house every book, and the librarian's route (EP dispatch) must avoid crowds.

Unlike a library, the "books" here were co-trained with the routing policy: you can't bolt on new expert shelves after the fact and expect the router to know they exist.
Compiler with specialized code paths

Think of each expert as a specialized kernel and the router as a JIT dispatcher: per input, a few kernels are selected and fused (weighted sum), while the rest stay cold in memory. Fine-grained MoE = many small kernels + rich dispatch; capacity factor = guard bands in your thread pool scheduling.

A JIT can recompile on profile data between runs; MoE routing is frozen at training time — deployment-time rebalancing only nudges selection bias, it doesn't retrain specialists.

09 Common Misconceptions

"A 671B MoE is 671B worth of quality in every token." No token ever touches more than ~37B of weights. Capacity is stored knowledge available across the routing space; per-token depth is the active budget. Quality gains are real but bounded by routing precision.

"MoE is 8× cheaper to serve than a dense 671B." It's ~18× cheaper in bandwidth and FLOPs per token, but you still buy ~0.7 TB of HBM per replica and pay all-to-all communication plus balancing overhead. Cheaper per token, not cheap to host.

"More experts is automatically better." More, smaller experts only help if the router learns to compose them; they add dispatch overhead and make load balancing harder (256 experts × top-8 is a much harder scheduling problem than 8 × top-2). The win came from fine-grained structure plus shared-expert isolation plus better balancing — together.

"Dropped tokens (capacity overflow) are a rounding error." At small batch sizes overflow is rare; at large EP deployments a bad α or a hot expert can silently drop a meaningful token fraction — quality loss that never shows up as an error, only as slightly worse outputs.

🔗
Connecting the dots: MoE attacks the weights side of the memory wall (doc 08) — read fewer weight bytes per token while storing more. The KV cache attacks the activations side: DeepSeek-V3's MLA is precisely that trick applied to attention state (doc 22). Both are answers to the same question: how do we keep growing capability while the wall stays fixed? Where the transformer itself might get replaced is the landscape of doc 25. And if the routing intuition feels fuzzy, rebuild it from first principles in doc 02.